Papers with real-world problems
Dive into Deep Learning for Natural Language Processing (D19-2)
Copied to clipboard
| Challenge: | GluonNLP is a powerful new toolkit that automates the most laborious aspects of deep learning for NLP. |
| Approach: | This hands-on tutorial demonstrates how to scale unsupervised pre-training techniques with Apache MXNet and GluonNLP. |
| Outcome: | This hands-on tutorial examines the challenges of scaling these models and algorithms effectively with Apache MXNet and GluonNLP. |
AutoNLU: An On-demand Cloud-based Natural Language Understanding System for Enterprises (2020.aacl-demo)
Copied to clipboard
| Challenge: | AutoNLU is an on-demand cloud-based system that enables users to create and edit datasets and train and test different state-of-the-art NLU models. |
| Approach: | They introduce an on-demand cloud-based system that provides an easy-to-use interface . they build powerful keyphrase extraction models that achieve state-of-the-art results . |
| Outcome: | The proposed model achieves state-of-the-art on two public benchmarks and is easy to use and use. |
NLP for Counterspeech against Hate and Misinformation (CSHAM) (2025.acl-tutorials)
Copied to clipboard
| Challenge: | tutorial aims to show how counterspeech is used to tackle abuse and misinformation by individuals, activists and organisations. |
| Approach: | tutorial aims to show how counterspeech is currently used to tackle abuse and misinformation . will also show how Natural Language Processing (NLP) and Generation (NLG) can be applied to automate its production. |
| Outcome: | The tutorial will bring diverse multidisciplinary perspectives to safety research . case studies from industry and public policy will be included . |
Investigating Prior Knowledge for Challenging Chinese Machine Reading Comprehension (2020.tacl-1)
Copied to clipboard
| Challenge: | ''Language is, at best, a means of directing others to construct similar-thoughts from their own prior knowledge,'' says K. S. Adams and Bruce. |
| Approach: | They present a free-form multiple-choice Chinese machine reading Comprehension dataset (C3) containing 13,369 documents and their associated 19,577 multiple-CHOice free- form questions. |
| Outcome: | The proposed model outperforms human models on linguistic, domain-specific, and general world knowledge problems. |
Self-supervised Regularization for Text Classification (2021.tacl-1)
Copied to clipboard
| Challenge: | Text classification models are prone to overfitting when limited texts are available for training. |
| Approach: | They propose a data-dependent regularization approach based on self-supervised learning . they define auxiliary tasks on input data without using human-provided labels . |
| Outcome: | Experiments on 17 text classification datasets demonstrate the effectiveness of the proposed method. |
Domain Adaptation with Adversarial Training and Graph Embeddings (P18-1)
Copied to clipboard
| Challenge: | Existing models for deep neural networks can handle data distributions between source and target domains, but they must deal with data distribution drifts. |
| Approach: | They propose a model that leverages unlabeled and labeled data from a related domain to deal with distribution drifts. |
| Outcome: | The proposed model improves over baselines on two real-world disaster datasets. |
MARCO: Multi-Agent Real-time Chat Orchestration (2024.emnlp-industry)
Copied to clipboard
Anubhav Shrimal, Stanley Kanagaraj, Kriti Biswas, Swarnalatha Raghuraman, Anish Nediyanchath, Yi Zhang, Promod Yenigalla
| Challenge: | MARCO is a multi-agent real-time chat orchestration framework for automating workflows that require interactions with tools, reasoning, and human collaboration. |
| Approach: | They propose a multi-agent real-time chat orchestration framework for automating workflows using LLMs. |
| Outcome: | The proposed framework performs with 94.48% accuracy and 92.74% accuracy on restaurant and retail conversations datasets and 44.91% improved latency and 33.71% cost reduction in a production setting. |
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models have enabled the development of coding agents for real-world code generation. |
| Approach: | They propose a novel LLM-driven test case generator that analyzes codebases and dependencies to generate test cases for real-world Python projects. |
| Outcome: | The proposed framework improves the performance of SWE-Bench by analyzing codebases and dependencies. |
Evaluating Multilingual Sentence Representation Models in a Real Case Scenario (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study has shown that the infamous Protocols are actually plagiarized . a convoluted task with no standard benchmarks for paraphrase detection and sentence similarity is a problem . |
| Approach: | They evaluate sentence representation models on the paraphrase detection task . they use a forged text from the so-called "Protocols of the Elders of Zion" scholars have demonstrated that the first text plagiarizes from the second . |
| Outcome: | The proposed model is based on the forged “Protocols of the Elders of Zion” . the model is similar to the standard model but has some problems . |
Deep Bayesian Active Learning for Natural Language Processing: Results of a Large-Scale Empirical Study (D18-1)
Copied to clipboard
| Challenge: | Existing studies on Active Learning (AL) for natural language processing have limited data requirements. |
| Approach: | They propose a Bayesian active learning approach that reduces deep learning's data dependence by comparing models and acquisition functions. |
| Outcome: | The proposed approach outperforms i.i.d. baselines and is more efficient than other approaches. |
Optimizing Annotation Effort Using Active Learning Strategies: A Sentiment Analysis Case Study in Persian (2020.lrec-1)
Copied to clipboard
Seyed Arad Ashrafi Asli, Behnam Sabeti, Zahra Majdabadi, Preni Golazizian, Reza Fahmi, Omid Momenzadeh
| Challenge: | Existing deep learning approaches require huge amounts of data to be trained properly. |
| Approach: | They propose to use Persian as a model to choose the samples for annotation instead of labeling the whole dataset. |
| Outcome: | The proposed models achieve the baseline performance with a significantly lower amount of labeled data. |
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced mathematical reasoning, but they still struggle with out-of-distribution (OOD) issues. |
| Approach: | They propose a framework to evaluate the logical validity of reasoning steps . they retrieves semantically similar questions and steps for PRM as a warmup . |
| Outcome: | The proposed framework outperforms baseline models on multiple real-world datasets. |
PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text Modeling (2021.acl-long)
Copied to clipboard
| Challenge: | PhotoChat contains 12k dialogues, each of which is paired with a user photo that is shared during the conversation. |
| Approach: | They propose to use PhotoChat to facilitate research on image-text modeling by combining a photo-sharing intent prediction task and a picture retrieval task to retrieve the most relevant photo according to the dialogue context. |
| Outcome: | The proposed tasks achieve 10.4% recall@1 and 58.1% F1 scores, indicating that the proposed dataset presents interesting yet challenging real-world problems. |
Fine-grained Information Extraction from Biomedical Literature based on Knowledge-enriched Abstract Meaning Representation (2021.acl-long)
Copied to clipboard
| Challenge: | Compared with general natural language texts, sentences from scientific papers usually possess wider contexts between knowledge elements. |
| Approach: | They propose a novel biomedical Information Extraction model to extract scientific entities and events from English research papers using Abstract Meaning Representation (AMR) they construct a sentence-level knowledge graph from an external knowledge base and encode it to improve the model's understanding of complex scientific concepts. |
| Outcome: | The proposed model can extract scientific entities and events from scientific literature and improve its understanding of complex scientific concepts. |
Unleashing the Power of Language Models in Text-Attributed Graph (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on graph learning on text-attributed graphs have been limited by memory cost and underutilization of relationships between nodes and words. |
| Approach: | They propose a Node Representation Update Pre-training Architecture based on Co-modeling text and graph to learn representations of papers and words simultaneously. |
| Outcome: | The proposed model outperforms baselines on the ogbn-arxiv benchmark dataset. |
Can LLMs Reason in the Wild with Programs? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models have shown superior capability to solve reasoning problems with programs. |
| Approach: | They propose a task where an LLM is tasked to solve a reasoning problem of unknown type by identifying the sub-problems and their corresponding formalisms. |
| Outcome: | The proposed model can be fine tuned to achieve better performance on ambiguous and mixed scope problems. |
How much coffee was consumed during EMNLP 2019? Fermi Problems: A New Reasoning Challenge for AI (2021.emnlp-main)
Copied to clipboard
| Challenge: | a new reasoning challenge is proposed to help AI systems to solve real-world problems . Fermi Problems are questions whose answers can only be approximated because their computation is either impossible or impossible. |
| Approach: | They propose a new reasoning challenge, Fermi Problems, which asks questions whose answers can only be approximated because their computation is either impractical or impossible. |
| Outcome: | The proposed datasets show that even fine-tuned large-scale language models perform poorly on these datasets. |
GreenKGC: A Lightweight Knowledge Graph Completion Method (2023.acl-long)
Copied to clipboard
| Challenge: | Knowledge graph completion (KGC) aims to discover missing relationships in knowledge graphs (KGs). |
| Approach: | They propose a modularized knowledge graph completion solution that learns embeddings for entities and relations through a score function. |
| Outcome: | Experimental results show that GreenKGC outperforms SOTA methods in low dimensions and even better against high-dimensional models with a much smaller model size. |
Scaling Collaborative Effort with Agents (2026.findings-acl)
Copied to clipboard
Shannon Zejiang Shen, Valerie Chen, Ken Gu, Alexis Ross, Zixian Ma, Jillian Ross, Alex Gu, Chenglei Si, Wayne Chi, Andi Peng, Jocelyn J Shen, Ameet Talwalkar, Tongshuang Wu, David Sontag
| Challenge: | Current evaluations of agents focus on producing high-quality, final outputs in one shot, failing to account for the inherently iterative nature of many real-world problems. |
| Approach: | They propose a framework that captures how an agent’s utility grows with increasing user involvement. |
| Outcome: | The proposed framework captures how an agent’s utility grows with increasing user involvement, revealing a missing ingredient in agent design: the ability to sustain engagement and scaffold user understanding. |
ToolCPT: Improving Tool Utilization in LLM Agents via Continuous Pre-training (2026.findings-acl)
Copied to clipboard
| Challenge: | Current approaches to enhancing tool use for LLM-based agents focus on post-training fine-tuning or test-time context extension. |
| Approach: | They propose to enhance tool knowledge for LLM-based agents during continuous pre-training . they curate 5.1 million code artifacts from large-scale, high-quality code repositories . |
| Outcome: | The proposed model outperforms existing methods on out-of-distribution tools on multiple benchmarks. |
What Do NLP Researchers Believe? Results of the NLP Community Metasurvey (2023.acl-long)
Copied to clipboard
Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, Samuel R. Bowman
| Challenge: | Getting sociological beliefs wrong can slow research and lead to wasted effort, missed opportunities, and needless fights. |
| Approach: | They present the results of the NLP Community Metasurvey, run from May to June 2022. |
| Outcome: | The NLP community metasurvey elicited opinions on controversial issues from May to June 2022. |
gMBA: Expression Semantic Guided Mixed Boolean-Arithmetic Deobfuscation Using Transformer Architectures (2025.findings-acl)
Copied to clipboard
| Challenge: | Mixed Boolean-Arithmetic (MBA) obfuscation protects intellectual property by converting programs into complex forms that are difficult to analyze. |
| Approach: | They propose a mixed-boolean-arithmetic (MBA) obfuscation framework that transforms a Transformer-based neural encoder-decoder into a truth table that is an automatically constructed semantic representation of an expression's behavior. |
| Outcome: | The proposed framework improves performance and highlights the importance of internal semantic expressions in recovering obfuscated code to its original form. |
TPS-Bench: Evaluating AI Agents’ Tool Planning & Scheduling Abilities in Compounding Tasks (2026.acl-long)
Copied to clipboard
| Challenge: | Large language model (LLM) agents have demonstrated strong problem-solving competence across domains like research and coding. |
| Approach: | They propose to use a tool repository to analyze the ability of large language model agents to solve complex problems. |
| Outcome: | The proposed model outperforms open-source and closed-source models in task completion rate and efficiency. |
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning (2026.acl-long)
Copied to clipboard
Afra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta, Jaehwan Jeong, Anisha Gunjal, Tahseen Rabbani, Maria Mazzone, David Randolph IV, Mohammad Mahmoudi Meymand, Gurshaan Chattha, Paula Rodriguez, Diego A. Mares Buendia, Pavit Singh, Michael Liu, Subodh Chawla, Peter Cline, Lucy Ogaz, Ernesto Gabriel Hernández Montoya, Zihao Wang, Pavi Bhatter, Marcos Ayestaran, Bing Liu, Yunzhong He
| Challenge: | Frontier models often lack a view of performance on open-ended, economically consequential tasks in high-stakes professional domains where practical returns matter most. |
| Approach: | They introduce a professional reasoning benchmark that recruits 182 qualified professionals to contribute questions inspired by their workflows. |
| Outcome: | The proposed model outperforms other models in 114 countries and 47 US jurisdictions on hard subsets. |